‹ BackNewsmodel pricing

model pricing

DeepSeek V4.1 Flash trails Kimi and GLM in third-party test, but runs at about one-seventh the cost
B.AI to End Free Access to DeepSeek V4 Flash on Sept 3, Start Discounted Billing
Google
2026-08-16 16:02:49

Gemini 3.7 Flash review: big coding gains, weaker reasoning and writing still show

Google launched Gemini 3.7 Flash on August 13 and made it generally available in more than 160 countries on day one. According to Decrypt’s review, the model accepts up to 1 million input tokens, returns 64,000 output tokens, handles images, video, audio, and PDFs, and can use tools while operating a computer. Google’s own benchmark sheet says the model beats Claude Sonnet 5 and GPT-5.6 Terra in 11 of 18 tested categories, including 1,588 Elo on Code Arena’s web development board and 30.4% on AutomationBench, though Decrypt notes those figures come from Google’s methodology and should be treated as company claims rather than settled fact. Decrypt’s hands-on tests found the sharpest improvement in coding. Gemini 3.7 Flash generated a playable browser game on the first try in 2 minutes and 13 seconds, a major step up from Gemini 3.6 Flash, which Decrypt said could not produce a working file in a similar test after its July 21 release. Results were less convincing elsewhere. In creative writing, Decrypt said Gemini produced a tidy story but broke the central prompt rule, losing to a free community model, Qwopus3.5-27B-v3. In associative reasoning, logic, and advanced math, the review said Gemini often showed decent structure but failed on crucial task requirements, including a bridge puzzle and a polynomial problem it left unfinished. Decrypt’s conclusion: Gemini 3.7 Flash is a strong low-cost execution model inside Google’s ecosystem, but its creativity and reasoning remain uneven.

1550
Gemini 3.7 Flash review: big coding gains, weaker reasoning and writing still show
Google launches Gemini 3.7 Flash three weeks after 3.6, with lower pricing aimed at coding and agent work
DeepSeek rolls out V4 Pro official release, claiming near-Claude performance at roughly 1/46 the cost
Alibaba
2026-08-04 01:51:01

Alibaba unveils Qwen3.8-Max with self-reported benchmark lead and aggressive token pricing

Alibaba’s Qwen team has introduced Qwen3.8-Max, a new flagship model that the company says scored 86.1 on OSWorld-Verified, ahead of GPT-5.6 Sol Max at 83.2, Fable 5 at 85.0, and Gemini 3.1 Pro at 76.2. The release also included a broader slate of benchmark claims, such as 93.0 on PaperBench and 86.6 on TerminalBench 2.1, alongside positioning the model for long-running autonomous work rather than standard chatbot use. The model uses a mixture-of-experts architecture with 2.4 trillion total parameters and about 95 billion active during inference, built on the Qwen3.5 architecture with a 1 million-token context window. Alibaba also said Qwen3.8-Max is suited for extended coding tasks, desktop software operation, experiment reproduction, and industrial workflows that feed visual input back into a planning loop. Pricing appears to be a central part of the launch. According to QwenCloud pricing cited in the report, Qwen3.8-Max costs $2 per million input tokens and $6 per million output tokens overseas, bringing the combined total to $8 per million tokens. That is below one-third of Claude Opus 5’s combined $30 and below one-quarter of GPT-5.6 Sol standard mode at $35. Still, the benchmarks and capability demonstrations were all disclosed by Alibaba and have not been independently verified, while the company has yet to publish the licensing terms for the promised open-weight release next week on Hugging Face and ModelScope.

1970
Alibaba unveils Qwen3.8-Max with self-reported benchmark lead and aggressive token pricing
OpenAI cuts prices for parts of the GPT-5.6 lineup and rolls out a new API mode